Papers with multilingual vision-language models

2 papers
Stop Pre-Training: Adapt Visual-Language Models to Unseen Languages (2023.acl-short)

Copied to clipboard

Challenge: Existing studies have shown that the pre-training in English does not transfer well to other languages in a zero-shot setting.
Approach: They propose a simple yet efficient approach to adapt VLP to unseen languages using MPLM.
Outcome: The proposed approach outperforms state-of-the-art models without large parallel corpora across three tasks.
Quantifying the Gaps Between Translation and Native Perception in Training for Multimodal, Multilingual Retrieval (2024.emnlp-main)

Copied to clipboard

Challenge: Existing models that account for perceptual differences in image captions are limited to use in English . culture-based tasks such as recognition, detection, and image retrieval are hindered by relying on English supervision.
Approach: They propose and evaluate caption augmentation strategies to address these gaps . they use captions from german perception and captions that have been machine-translated or human-transcribed from English into german .
Outcome: The proposed models achieve a mean recall improvement of +1.3, but still lack flexibility . cultural differences present in language with respect to object specificity and importance .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations